Back

Pilot and Feasibility Studies

Springer Science and Business Media LLC

Preprints posted in the last 7 days, ranked by how well they match Pilot and Feasibility Studies's content profile, based on 14 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Acceptability, feasibility, quality of life and diabetes distress score outcomes: A pragmatic randomised clinical trial on continuous glucose monitoring for people with type 1 diabetes

Marban-Castro, E.; Muhwava, L.; Girdwood, S.; Kemp, T.; Freitas, J.; Kamau, Y.; Otieno, M.; Akach, D.; Morato, A.; Sanz, S.; Fiechter, V.; Erkosar, B.; Watson, M.; Vetter, B.; Haldane, C.; Shilton, S.; Rheeder, P.; Dave, J. A.; Carrihill, M.; Karsas, M.

2026-08-31 endocrinology 10.64898/2026.08.26.26361479 medRxiv
Top 0.2%
2.8%
Show abstract

Introduction: Continuous glucose monitoring (CGM) offers an advancement over traditional self-monitoring of blood glucose (SMBG) for people living with type 1 diabetes (T1D). However, evidence on the acceptability and feasibility of different CGM use cases in African populations remains limited. Methods: This was a pragmatic three-arm, randomised controlled trial on CGM conducted among people living with T1D in three public healthcare clinics in South Africa. Participants were assigned to Arm 1 (continuous CGM), Arm 2 (periodic CGM), or Arm 3 (SMBG). Diabetes education was provided at all study visits. Feasibility was assessed by adherence to CGM use and through the Glucose Monitoring Satisfaction Survey (GMSS). Diabetes distress was measured by the Diabetes Distress Scale (DDS), health-related quality of life (HRQoL) by the EQ-5D scales, and acceptability using the Theoretical Framework of Acceptability (TFA). Surveys were collected on paper and transferred to OpenClinica. Analyses were performed in R. The trial was registered in the Clinical Trials Registry (NCT05944718) on July 13, 2023. Results: A total of 83 participants were included in Arm 1, 85 in Arm 2, and 80 in Arm 3. CGM mean active time was 55% in Arm 1 versus 69% in Arm 2. The proportion of participants meeting the [≥]70% active time threshold was higher in Arm 2 (52%) than in Arm 1 (34%). Diabetes' distress declined across arms during the intervention period, with no significant difference between arms; distress increased slightly six months post-intervention but remained below baseline. At 6 months, glucose monitoring satisfaction was significantly higher in both CGM arms than in the SMBG arm, and satisfaction increased over time in CGM arms. Health-related quality of life remained stable across arms during the intervention period with no significant difference between arms. High acceptability was observed in both CGM arms, with higher ratings in the periodic arm. Conclusions: CGM was acceptable to people living with type 1 diabetes and feasible to use in public-sector clinics in South Africa, with high acceptability under continuous and periodic use. Health-related quality of life remained stable across arms, and diabetes-related distress declined, during the intervention period, across arms. Glucose monitoring satisfaction rose significantly in both CGM arms compared to SMBG. Periodic CGM might be a promising and potentially more scalable option than continuous use for public-sector care.

2
Usability, acceptability and feasibility of continuous glucose monitoring among children and adolescents with type 1 diabetes in Kenya

Amolo, P.; Mungai, L.; Karume, A. K.; Kibugi, J.; Mwende, W.; Botella, N.; Haldane, C.; Kamau, Y.; Marban-Castro, E.

2026-09-01 endocrinology 10.64898/2026.08.27.26361447 medRxiv
Top 0.2%
2.2%
Show abstract

Introduction Continuous Glucose Monitoring (CGM) is considered standard care in high-income countries. There is, however, limited published evidence on CGM use in low- and middle-income countries. The purpose of this study was to assess the usability, acceptability, and feasibility of CGM use among people living with type 1 diabetes (T1D) and caregivers in a low-resource setting. Research Design and Methods This prospective study conducted at the Kenyatta National Hospital purposively enrolled persons aged 4-25 years who had been on management for T1D for at least six months, and caregivers of those under 18 years. Fourty youth living with T1D used CGM for three months in place of self monitoring of blood glucose (SMBG). The System Usability Scale (SUS), a Theoretical Framework of Acceptability-based questionnaire, the Diabetes Distress Scale (DDS), the Glucose Monitoring Satisfaction Survey (GMSS), and a feasibility survey were administered. Outcomes were summarized descriptively, including means, medians, and frequencies using R statistical software. Results The median SUS score was 98.8 (IQR 92.5-100.0). Acceptability was high, and the median total GMSS score improved from 3.73 to 4.73. Among adolescents and adults, the median overall DDS score reduced from 1.54 to 1.36, with reductions in scores in all domains, except for hypoglycemia distress which increased, and physician distress which remained low. Among caregivers, the median overall DDS score declined from 2.05 (moderate distress) to 1.90 (low distress), with modest reductions in teen management and parent-teen relationship distress and a slight increase in personal distress. Median CGM active wear time was 89%. Conclusion This study comprehensively evaluated CGM across usability, acceptability, and feasibility outcomes, with the findings supporting the integration of CGM into routine diabetes management in low-resource settings. The short follow-up period, however, may not capture changing perceptions or long-term adherence.

3
GLP-1/GIP Uptake, Indication, and Access Pathways Among US Adults in the Understanding America Study

Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.

2026-09-02 endocrinology 10.64898/2026.08.28.26361368 medRxiv
Top 0.3%
1.7%
Show abstract

Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.

4
Sensorimotor effects of heatwrap and exercise in acute low back pain: results of a randomised controlled trial

Cote Picard, C.; Desgagnes, A.; Tittley, J.; Mailloux, C.; Perreault, K.; Mercier, C.; Dionne, C. E.; Roy, J.-S.; Masse-Alarie, H.

2026-09-02 rehabilitation medicine and physical therapy 10.64898/2026.08.31.26361841 medRxiv
Top 0.3%
1.7%
Show abstract

Background: Heatwrap is recommended for acute low back pain (ALBP), and previous research found heatwrap plus exercise more effective than each intervention alone. While recommended by clinical guidelines, their impact on mechanistic outcomes is unknown. This trial aimed to (i) assess immediate and short-term effects of heatwrap alone or combined with exercise, compared with a sham heatwrap, on pain sensitivity, lumbar muscle activity, current pain intensity, and trunk flexion range of motion, and (ii) explore whether changes in pain sensitivity and lumbar muscle activity are associated with changes in clinical symptoms from baseline to 1-week follow-up. Methods: A randomised controlled trial took place at a single research center. Of 315 individuals screened for eligibility, 99 adults with ALBP were recruited and assigned to one of three intervention groups: heatwrap plus exercise (n=34), heatwrap alone (n=33) or sham heatwrap (n=32). Interventions were applied for one hour at the first visit, and immediate effects were measured. Then, interventions were applied for 7 days, and short-term effects were measured at 1-week follow-up. Outcomes included pressure pain threshold, temporal summation of pain, flexion-relaxation ratio, trunk range of motion and current pain intensity. Results: Heatwrap and exercise did not produce greater effects over time than heatwrap alone or a sham heatwrap on all outcomes, and changes in sensorimotor outcomes at one week were not associated with changes in symptoms. Conclusions: Heatwrap and/or exercises did not influence specifically the potential sensorimotor mechanisms tested in individuals with ALBP. Trial registration: ClinicalTrials.gov; registration number: NCT03986047

5
Spatial Geometry and Prevalence of Tunneling and Undermining in Pressure Ulcers

Frade, S.; Tunyiswa, Z.; Shin, M.; Dirks, R.

2026-09-01 dermatology 10.64898/2026.08.28.26361615 medRxiv
Top 0.4%
1.2%
Show abstract

Background: Pressure ulcers often develop complex three-dimensional morphologies that extend beyond the visible wound surface. Subsurface extensions such as tunneling and undermining create hidden cavities that complicate clinical assessment and wound management. Despite their clinical relevance, the prevalence and spatial characteristics of these subsurface wound morphologies have not been well characterized at scale. Methods: We performed a registry-based analysis using data from the LIFT-OFF Pressure Ulcer Registry, which captures longitudinal clinical documentation of pressure ulcers treated in routine care. The registry included approximately 18,000 patients with 32,000 documented pressure ulcers. Spatial characteristics of tunneling and undermining were analyzed using measurements recorded during routine wound assessments, including tract length, direction, and circumferential extent. Directional and circumferential distributions of subsurface defects were examined to characterize wound geometry. Results: Tunneling was present in 764 of 14,700 full-thickness pressure ulcers (5.2%), whereas undermining occurred in 2,293 wounds (15.6%). Tunneling tracts were typically short and exhibited directional clustering relative to the wound bed. In contrast, undermining demonstrated broader circumferential distributions and frequently involved larger subsurface separations beneath the wound margin. Both morphologies demonstrated distinct spatial patterns across anatomical locations and wound stages. Conclusion: Tunneling and undermining are common subsurface features of pressure ulcers and exhibit distinct spatial geometries. Whereas tunneling manifests as directional tract-like extensions, undermining more frequently produces circumferential tissue separation beneath wound margins. Improved characterization of subsurface wound architecture may enhance assessment of wound complexity and provide information not captured by surface measurements alone. Future studies should evaluate whether these features contribute to wound severity assessment, prognosis, and risk stratification.

6
Effects of collaborative clinical visit agenda-setting interventions: A systematic review and meta-analysis

Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.

2026-09-03 medical education 10.64898/2026.08.30.26361729 medRxiv
Top 0.5%
1.2%
Show abstract

Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.

7
Electronic health data exploring cardiorespiratory responses of transfusions in preterm infants: An international multicenter cohort study

Honore, A.; Rech, T.; Scrivens, A.; Binotto, I.; Zandvoort, C. S.; van der Staaij, H.; Peck, M.; Zivanovic, S.; Stanworth, S. J.; Hartley, C.; Dame, C.; Deschmann, E.; the Neonatal Transfusion Network,

2026-09-03 pediatrics 10.64898/2026.09.01.26361418 medRxiv
Top 0.5%
1.2%
Show abstract

Background and Objectives: Preterm infants are commonly transfused, yet direct cardiorespiratory effects of red blood cell (RBC) transfusions remain poorly understood. We explored the feasibility of using multicentre electronic health data (EHD) to study such cardiorespiratory responses. Methods: Highly granular routine EHD were collected from preterm infants born <32 weeks gestational age at three European centres. Heart rate, oxygen saturation, and respiratory rate were evaluated 12 hours before and after the RBC transfusion. Results: A total of 321 transfusions in 164 infants were analysed. Overall, there was no significant change in the rate of bradycardia and apnoea following transfusion. Cardiorespiratory parameters varied substantially between infants; e.g. 20% of transfusions were associated with an unexpected, significant increase in heart rate. Respiratory rate and oxygen saturation exhibited similarly heterogenous patterns following transfusion. In sub-group analysis, the proportion of transfusions with increased heart rate was significantly higher within the first two weeks than later (32% vs 13%, p=0.0019). Conclusions: Multicentre EHD extraction allows to identify otherwise masked short-term effects of RBC transfusions on cardiorespiratory parameters, possibly indicating cardiac or pulmonary overload. Such effects may vary with adaptation to anaemia. Analysing EHD may ultimately enable personalized transfusion practice.

8
A Measurement-Based Care Strategy for Buprenorphine-Naloxone Treatment (Bup-MBC): Development of an EHR-Integrated Intervention

Reese, T.; Audet, C.; Ancker, J.; Wright, A.; Marcovitz, D.; Kast, K. A.; Bridges, J.; Tindle, H.; Shah, M.; von Horn, A.; Matheny, M. E.

2026-09-01 addiction medicine 10.64898/2026.08.27.26361539 medRxiv
Top 0.5%
1.1%
Show abstract

Introduction: Risk of recurrent opioid use during buprenorphine-naloxone (bup-nx) treatment is dynamic and remains elevated after initiation, with vulnerability shaped in part by treatment intensity and gaps between visits, yet routine outpatient care relies on episodic encounters and retrospective data. This mismatch can delay recognition of emerging instability and limit timely treatment adjustments. This paper reports the development and specification of an intervention strategy to address this mismatch. Methods: We used a structured, multi-phase design process to specify and configure a measurement-based care (MBC) strategy for bup-nx treatment (Bup-MBC) in outpatient addiction clinics through three phases: (1) a systematic review of patient-reported outcome measures (PROMs) for substance use treatment; (2) a qualitative needs assessment using the Theoretical Domains Framework and COM-B (Capability, Opportunity, Motivation-Behavior) model to identify gaps in risk monitoring, agency, and trust; and (3) iterative co-design with multidisciplinary clinicians to refine workflow fit and trust-preserving use of data. Patients informed item and feedback content during the needs assessment but did not participate in the co-design cycles. Results: Bup-MBC integrates (1) brief between-visit PROMs (e.g., withdrawal, craving, adherence); (2) immediate non-punitive patient feedback; (3) clinician-facing summaries and non-directive prompts in the electronic health record (EHR); and (4) an opt-in between-visit outreach pathway with predefined safety triggers, all configured within existing EHR and patient portal infrastructure. It targets patient and clinician capability to recognize changes in risk, opportunity for action through structured monitoring and visit preparation, and trust and agency through non-punitive communication, without adding substantial burden. The full measure set, severity bands, and question-to-action map are provided as supplementary material. Key trade-offs included prioritizing single-item measures for feasibility, balancing opt-in outreach with safety overrides, and assuming routine clinician use of summaries. Conclusion: This development study specifies an EHR-integrated MBC strategy for outpatient bup-nx treatment. As single-center design work with co-design limited to clinicians and delivery contingent on portal or text-message access, its outputs are hypotheses about mechanism and fit rather than demonstrated effects. Feasibility studies are needed to evaluate uptake, acceptability, workflow fit, and effects on treatment.

9
The effectiveness of a complex intervention, aimed at reducing hospital occupancy, to improve Emergency Department patient flow: a retrospective controlled interrupted time series

McHenry, R. D.; Caesar, D.; Clarke, B.; Mackay, D.; Pell, J.

2026-09-03 health systems and quality improvement 10.64898/2026.08.31.26361802 medRxiv
Top 0.6%
1.1%
Show abstract

Objectives Emergency department (ED) crowding is recognised as an important public health concern internationally, and is driven principally by exit block, the shortage of inpatient beds for patients requiring admission. This study aimed to evaluate whether a complex intervention targeting hospital occupancy improved ED patient flow, and quantified the change in attendances. Methods A controlled interrupted time series using weekly, publicly reported Public Health Scotland data from 1 January 2022 to 1 February 2026. The multi-component intervention focused on reducing hospital occupancy and included additional adult social care funding; engagement with regional social care providers; accelerated implementation of the Discharge without Delay programme; re-evaluation of whole-hospital escalation thresholds and response; resource and data supporting inpatient department reductions in length of stay; and additional investment in remote clinical assessment. The intervention commenced at a large tertiary ED on 01 February 2025. Primary outcomes were the proportions of attendances spending [&ge;]4, [&ge;]8 and [&ge;]12 hours in the ED. The secondary outcome was attendance volume. Segmented regression was fitted with a contemporaneous control series, seasonal terms and autoregressive moving average errors. Long waits were additionally illustrated as potentially avoided deaths. Results The analysis covered 161 pre-intervention and 52 post-intervention weeks. Relative to pre-intervention levels, the proportion of attendances waiting over 4 hours fell by 10.4% (95% CI 1.6 to 19.2%), by 16.4% (95%CI 1.3 to 31.5%) over 8 hours and by 24.3% (95%CI 2.6 to 46.1%) over 12 hours. Using established associations between long ED waits and excess mortality, by one-year the intervention was potentially associated with 54 fewer excess deaths (95%CI 19 to 93). Attendances rose by 3.8% (95%CI 1.3 to 6.4%) against the counterfactual. Conclusions A complex intervention targeting hospital occupancy was associated with a reduction in long ED waits despite rising attendances. Interventions addressing hospital occupancy can meaningfully improve ED crowding.

10
The use of computerised testing to assess cognitive performance in people with HIV in South Africa

Edmond, E. C.; Dreyer, A. J.; Winston, A.; Khoo, S. H.; Joska, J.; Nightingale, S.

2026-08-31 hiv aids 10.64898/2026.08.27.26361083 medRxiv
Top 0.7%
0.8%
Show abstract

Background Computerised cognitive testing may address the global challenge in identifying cognitive changes in people living with HIV scalably and affordably. We assessed a computerised battery (CB) of cognitive tests, in a prospective cohort (CONNECT) of people with HIV in a low-income peri-urban area of Cape Town, South Africa during a national programmatic switch from efavirenz- to dolutegravir-based antiretroviral therapy (ART). Methods We recruited 170 people with HIV and 91 people without HIV (controls) (140[82%] and 41[45%] followed up). The CB and gold-standard pen&paper cognitive testing (P&P) were performed at both timepoints. Technology familiarity/use questionnaire data were also collected. We compared performance in detecting lower group-level cognitive performance associated with efavirenz treatment. Furthermore, the CB was compared to P&P in classifying individuals with low cognitive performance, correlation of global test scores and domain-level scores between batteries, and practice effects between timepoints. Exploratory principal component analysis was also performed. Results People with HIV on efavirenz at baseline had lower performance on the computerised battery than controls, {Delta}T=2.6, p=0.0047. This difference was lost after switching to dolutegravir-based ART at follow-up. CB and P&P global T were moderately correlated (R2=0.203, p<0.001), and the CB performed moderately in classification of low cognitive performance against the gold standard (AUC 0.70, sensitivity 0.52, specificity 0.76, PPV 0.40, and NPV 0.84). Selecting the first three principal components improved both classification of low cognitive performance (AUC 0.77) and correlation strength with P&P global T (R2=0.3, p<0.001). The CB did not show practice effects. Most participants owned a mobile phone (95%, 85.9% of these smartphones). Performance was better in smartphone owners ({Delta}T=1.8) and computer owners (23%, {Delta}T=1.8). Conclusions Delivering computerised cognitive testing was feasible in this low-income southern African setting. The CB showed reasonable construct validity (detecting known lower cognitive performance associated with efavirenz-ART) and may detect broad cognitive characteristics such as processing speed and accuracy. However, correlation of CB results with gold standard P&P testing was low-moderate and may limit its applicability as a diagnostic tool. This might be improved by including a wider range of cognitive domains tested in the CB, or data driven analysis. Brief CBs may fulfil an initial screening role, followed by more detailed clinical assessment.

11
Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation

Chowdhury, A. R.; Chowdhury, B.

2026-09-02 health informatics 10.64898/2026.09.01.26361908 medRxiv
Top 0.7%
0.8%
Show abstract

Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.

12
Characterizing Preschool Children's Multi-Matrix Air Pollutant Exposures Across Home and Early Childhood Education Settings: A Paired Silicone Wristband Study Protocol

Mutic, A. D.; McCauley, L.; Andrew, A.; Fitzpatrick, A.

2026-08-31 occupational and environmental health 10.64898/2026.08.27.26361470 medRxiv
Top 0.9%
0.6%
Show abstract

Background: Children spend more than 90% of their time indoors, and early childhood education settings (ECEs) are an understudied, high-occupant-density indoor microenvironment where exposure to volatile organic compounds, particulate matter, and other toxicants has been documented. Limited knowledge exists on ECE-specific exposures affecting young children and how they compare to exposures in the home. Methods: This prospective, repeated-measures pilot study targeted enrollment of 44 preschool-aged children and 8 ECE staff across two geographically and sociodemographically distinct ECEs in metropolitan Atlanta, Georgia. Paired silicone wristbands, one home-designated and one ECE-designated, were exchanged between settings across three consecutive days and nights beginning at enrollment to characterize microenvironment-specific exposure. A single spot urine sample was also collected from each child. Continuous indoor air quality monitoring was conducted in two classrooms per site. Caregivers and ECE staff completed structured questionnaires assessing home and ECE environmental characteristics, child respiratory risk, and protocol feasibility and acceptability. Feasibility was evaluated using eight pre-specified indicators spanning recruitment and enrollment, wristband wear duration and loss by microenvironment, urine sample collection completeness, and survey completion by instrument and respondent group. Conclusion: This pilot will establish feasibility and acceptability parameters for a paired, multi-matrix silicone wristband protocol across home and ECE microenvironments. Findings will inform the design, sample size, and power calculations for a subsequent study testing indoor air interventions and pediatric respiratory outcomes in ECEs. Feasibility outcomes are reported in a companion manuscript.

13
Household hunger trajectories and mental health symptoms of adolescent girls and young women in rural South Africa: an HPTN-068 longitudinal study

Chakraborty, R.; Rosenberg, M.; Weigel, M. M.; Pettifor, A.; Kahn, K.; Gomez-Olive, F. X.

2026-09-03 public and global health 10.64898/2026.09.01.26361988 medRxiv
Top 0.9%
0.6%
Show abstract

Purpose Despite the high documented prevalence of hunger and poor mental health in adolescent girls and young women (AGYW) in South Africa, this relationship remains understudied, with existing studies limited by their cross-sectional designs. This longitudinal study aimed to identify the association of hunger trajectories with anxiety and depressive symptoms, and hope in AGYW. Methods We used secondary data from the HIV Prevention Trials Network (HPTN) -068 conducted in rural Agincourt, South Africa. Complete data from 1779 AGYW collected at baseline (2011/12) and three annual follow-up visits were used. Hunger trajectories, measured using the Household Hunger Scale, were estimated via Group-Based Trajectory Modelling. Self-reported incident anxiety and depressive symptoms and hope were assessed based on AGYWs last two follow-up visits. Covariate adjusted modified Poisson regression models estimated the association between hunger trajectories and incident anxiety symptoms, incident depressive symptoms, and hope. Results Moderate-severe hunger was prevalent in 11.0%, 10.8%, and 6.0% of the households at baseline, follow-up 1, and 2, respectively. Incident anxiety symptoms were reported by 4.5%, incident depressive symptoms by 20.0% and hopelessness by 52.8% of the AGYW. Two hunger trajectories were identified- no hunger (82%) and marginal hunger (18%). Hunger trajectories were not associated with incident anxiety symptoms [RR:1.09, 95% CI: 0.55, 2.18], incident depressive symptoms [RR: 0.97; 95% CI: 0.72, 1.33] nor hope [RR: 1.00; 95% CI: 0.81, 1.23] in AGYW. Conclusion Better understanding of the factors that promote resiliency and mental health of AGYW in this setting is warranted to inform the design of interventions.

14
Patient factors influencing participation in aerobic exercise during inpatient and outpatient rehabilitation post stroke: a prospective cohort study

Barzideh, A.; Devasahayam, A. J.; Marzolini, S.; Munce, S.; Sibley, K. M.; Inness, E. L.; Mansfield, A.

2026-08-31 rehabilitation medicine and physical therapy 10.64898/2026.08.26.26361451 medRxiv
Top 0.9%
0.5%
Show abstract

Background: Aerobic exercise is recommended during stroke rehabilitation to improve cardiorespiratory fitness and support recovery; however, participation rates remain low. While institutional and system-level barriers have been widely examined, less is known about how individual patient factors influence engagement in aerobic exercise during rehabilitation. Objectives: We aimed to determine whether depressive symptoms, apathy, self-efficacy and outcome expectations for exercise, perceived barriers, or past exercise history were associated with aerobic exercise participation in stroke rehabilitation. Methods: In this prospective cohort sub-study, adults admitted to in- or out-patient stroke rehabilitation at three urban hospitals completed validated questionnaires assessing depressive symptoms, apathy, exercise self-efficacy, outcome expectations for exercise, perceived barriers to being active, and premorbid exercise history. Participants were separated into two groups for analysis: those who completed aerobic exercise during rehabilitation and those who did not. Equivalence testing and between-group comparisons were performed. Results: Sixty-two participants were enrolled; 16 participated in aerobic exercise and 46 did not. Groups were not equivalent on any individual-level factors. Compared to non-participants, those who performed aerobic exercise had significantly higher depressive symptom scores (p=0.0025) and lower self-efficacy for exercise (p=0.0087). Non-participants demonstrated significantly higher apathy (p=0.0007). No significant differences were found for outcome expectations, perceived barriers, or exercise history. Conclusion: Depressive symptoms and lower self-efficacy did not impede aerobic exercise participation during rehabilitation. Increased apathy, however, was associated with non-participation. Findings highlight the need for individually tailored aerobic exercise prescriptions that consider motivational and affective factors to optimize engagement during stroke rehabilitation.

15
Half of alcohol, drug, and self-harm presentations cannot be identified in coded emergency department data: a diagnostic accuracy study of a large language model

Humphries, C.; Brett, J.; Gruber, F.; James, E.; McKendrick, T. I.; McNairn, K. C.; Miell, A.; O'Brien, R.; Rahman, F.; Schölin, L.; Stewart, M.; Casey, A.

2026-08-31 health informatics 10.64898/2026.08.26.26361443 medRxiv
Top 1.0%
0.5%
Show abstract

Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients [&ge;]16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review.

16
What Matters Most: A Multi-Stakeholder Study of Outcome Domains in Lower-Limb Prosthesis Use

Ahmed, M. E.; Karlsson-Brown, S.; Koufaki, P.; Ahmadi, M.; Mico-Amigo, E. M.

2026-09-03 rehabilitation medicine and physical therapy 10.64898/2026.08.31.26361544 medRxiv
Top 1%
0.5%
Show abstract

Purpose: Lower-limb prosthesis use involves interacting physical, psychosocial, and device-related outcomes that may not be fully captured by conventional clinical assessment. This study aimed to develop and evaluate a stakeholder-informed framework of outcome domains relevant to meaningful everyday prosthesis use. Materials and Methods: A mixed-methods participatory design comprised a structured synthesis of selected clinically relevant content from five established patient-reported outcome measures; semi-structured interviews and importance and actionability ratings with 18 contributors (12 prosthesis users, four clinicians, and two industrial partners); and integration of the synthesis, qualitative, and rating findings. Interview records were analysed using reflexive thematic analysis, and ratings were analysed descriptively. Results: The resulting framework comprised four interrelated domains: Mobility, Physical Function, Psychosocial Wellbeing, and Prosthesis Experience. Mobility showed the clearest convergence across stakeholder perspectives. Prosthesis users showed the largest importance actionability gap for Prosthesis Experience (4.5 vs 3.0), whereas clinicians showed the largest gap for Psychosocial Wellbeing (5.0 vs 3.0). Interviews highlighted day-to-day variability in prosthesis use and the influence of confidence, fatigue, comfort, environmental conditions, social context, and device usability. Conclusions: Meaningful outcome assessment in prosthetic rehabilitation should extend beyond mobility alone to consider physical function, psychosocial wellbeing, and prosthesis experience within everyday contexts. The proposed framework provides a stakeholder-informed foundation for multidimensional outcome assessment in prosthetic rehabilitation.

17
CHARMS and PROBAST+AI: an updated template for Data Extraction and Risk of Bias Assessment in systematic reviews of prediction models

Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.

2026-08-31 cardiovascular medicine 10.64898/2026.08.26.26361189 medRxiv
Top 1%
0.5%
Show abstract

Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.

18
A Pragmatic Randomized Trial of an EHR-Integrated Generative AI Chart Summarization Tool for Ambulatory Clinicians

Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.

2026-08-31 health informatics 10.64898/2026.08.26.26361496 medRxiv
Top 1%
0.5%
Show abstract

BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.

19
People living with multiple long-term conditions have different pathways of unscheduled care in hospital: findings from an analysis of routinely-collected clinical data

Witham, M.; Evison, F.; Bellass, S.; Cooper, R.; Gallier, S.; Pretorius, S.; Sapey, E.; Suklan, J.; Sayer, A. A.

2026-09-01 health informatics 10.64898/2026.08.28.26361696 medRxiv
Top 1%
0.4%
Show abstract

Study Objective Little is known about where in hospital care for multiple long-term conditions (MLTC) is delivered. We aimed to describe pathways of care (ward transfers) and outcomes for people admitted to hospital for unscheduled care by MLTC status and other key sociodemographic characteristics. Design and setting Analysis of routinely-collected electronic health records from a large acute UK hospital. Participants Adult unscheduled care admissions from 1st July 2018 to 30th June 2019. The presence of two or more of 59 long-term conditions was ascertained using ICD-10 codes from previous hospital discharges. Main outcome measures Markov state transition probabilities were derived for ward moves and compared for MLTC vs no MLTC, age, sex, ethnicity and neighbourhood deprivation. Outcomes (length of stay, death, readmission, move from definitive ward) and time spent in emergency and assessment departments were compared between subgroups. Results A total of 33,252 adults, mean age 56.0 (SD 21.9) years were analysed; 14,834 (42.4%) had MLTC. People with MLTC were more likely to die in hospital (4.2 vs 1.9%, p<0.001), transfer to internal medicine wards or older peoples medicine wards, were less likely to transfer to surgical wards, had longer median length of stay (1.83 vs 0.69 days, p<0.001), stayed longer in acute medical units (15.5 vs 9.6 hours, p<0.001), and were more likely to move from their definitive ward (18.2 vs 16.4%, p=0.002). Conclusion Unscheduled hospital care pathways are complex and differ for people with MLTC, who have worse outcomes and may be less likely to receive optimal care.

20
Sleep Intervention for NHS Healthcare Shift Workers: A Pilot Study of Noise-Masking Earbuds

Hickman, R.; Joyce, D. W.; Gray, N.; Shergill, S.; D'Oliveira, T. C.

2026-09-01 psychiatry and clinical psychology 10.64898/2026.08.28.26361673 medRxiv
Top 1%
0.3%
Show abstract

Background: Shiftwork disrupts natural sleep-wake cycles, alters light exposure patterns, and contributes to circadian misalignment. Detrimental health consequences associated with shift work include elevated risk for metabolic disorders, cardiovascular disease, cancer and all-cause mortality. Healthcare workers have one of the highest rates of shift work exposure, yet there are relatively few non-pharmacological interventions (with good evidence) developed to improve sleep outcomes in this population. Objective: A pre-post pilot interventional study assessed the acceptability and perceived effectiveness of commercial noise-masking earbuds on improving subjective sleep characteristics among National Health Service (NHS) healthcare staff working fast rotating shifts. Methods: Noise-masking sleep earbuds (Kokoon NightBuds) were worn for a pilot six-week intervention by twenty-seven NHS nurses (aged 26-43 years, 88.9% female) working fast rotating shifts from the EClocker Study. Sensors inside the earbuds were paired with a smartphone app to monitor sleep. An audio library in the smartphone app delivered personalised relaxation exercises and sleep techniques drawn from cognitive behavioural therapy for insomnia (CBT-I). A pre-post two-week monitoring period with daily smartphone-based Experience Sampling Methods (ESM) captured perceived daily sleep patterns. Acceptability and perceived effectiveness of the earbuds in promoting better sleep outcomes was assessed. Results: Use of the noise-masking sleep earbuds over a six-week period was associated with positive sleep improvement trends and elicited promising acceptability. Almost two thirds of NHS fast rotating shift nurses (63%) subjectively reported reductions in general sleep disturbance symptoms (PSQI Global), one in four experienced perceived sleep quality improvements (SQ; 25.9%), one in five reported sleeping longer (TST; 22.2%), and a third perceived falling asleep faster (SOL; 33.3%), had better sleep efficiency (SE; 33.3%) and improved daytime dysfunction (33.3%) (PSQI subcomponent scores). Sleep diaries (CSD) collected daily using smartphone-based ESM also demonstrated small improvements post-sleep earbud use; nurses reported sleeping an average 18 minutes longer (TST) and fell asleep more easily, on average 11 minutes faster (SOL). Sleep earbuds were generally well tolerated; 56% of nurses reported the earbuds as (somewhat to very) helpful, 52% reported (somewhat to strongly) falling asleep more easily (SOL), 44% felt (somewhat to strongly) their sleep quality was improved (SQ) and 30% agreed (somewhat to strongly) they slept longer (TST) and had less disturbed sleep. Conclusions: To our knowledge, this is the first study in Europe to pilot noise-masking earbuds as a potential non-pharmacological aid to improve sleep-wake behaviours or mitigate fatigue for healthcare staff. Preliminary results showed promising acceptability and (small) perceived sleep improvement trends following a targeted six-week earbud intervention in NHS fast rotating shift nurses.